ci: [branch-1.1] run every tier on release-branch pull requests, and the full suite before an RC (#6218) - #6284
Conversation
…e before an RC (apache#6218) * docs: document CI coverage on release branches and run the full suite before an RC Release branches get only the PR tier: GitHub fires schedules only on the default branch, the merge queue covers only main, and ci.yml runs on push only for main. Document that, add a step that dispatches ci.yml and miri.yml on the release branch before tagging each release candidate, and correct the stale claim that the release branch runs the same CI as main. Also keep the docs deploy to main. A dispatch runs every job, so dispatching ci.yml on a release branch would have published that branch's docs over the website. * ci: run every tier on release-branch pull requests A release branch has no merge queue and no nightly: the queue ruleset targets the default branch, ci.yml runs on push only for main, and GitHub fires schedule only on the default branch. So a tier that holds a suite back for the queue or the nightly drops it there instead. A pull request against branch-N.M now runs every pr, queue and nightly job, still routed by the path filters. docs stays push-only and Spark 3.4 stays label-only, so a labeled run there adds Spark 3.4 or nothing. The base branch reaches compute-changes.py as PR_BASE_REF. check-ci-config.py now pins every variable event_from_env reads to the github context expression ci.yml sets it from, because a dropped or mistyped one reads as an empty string and changes the routing without failing anything. * ci: pin the docs deploy guard, and check it before dispatching a release branch A dispatch runs the release branch's own ci.yml, docs job included, so the promise that it leaves the website alone holds only on a branch that carries the refs/heads/main guard. branch-1.1 was cut without it. Both dispatch instructions now start by printing the branch's docs job, and check-ci-config.py fails if the guard is dropped from ci.yml. The dispatch result check now queries each workflow on its own. With --limit 2 across both, a second ci.yml dispatch could push the Miri run out of the list. Also correct two statements: the documentation and change log pull requests change only Markdown, so they run no heavy suite, and a CI change on main reaches a release branch only if it is backported. (cherry picked from commit 2d1ee49)
sunchao
left a comment
There was a problem hiding this comment.
Summary
- Prior state and problem: Release-branch PRs ran only the PR tier, leaving queue and nightly coverage absent. Manually dispatching CI on a release branch could also publish its documentation over the website.
- Design approach: Recognize
branch-N.Mtargets, enable all three existing tiers subject to path filters, and restrict site deployment tomain. - Correctness / compatibility analysis: Routing tests confirmed the intended release-branch behavior and unchanged non-release behavior. Existing Spark profiles remain intact, including the Spark 3.4 SQL suite's label requirement. Spark execution semantics are untouched.
- Key design decisions: Reusing
POLICYand a small release-branch helper avoids duplicating job matrices. Additional compute is confined to the intentionally expanded CI coverage. Label events avoid repeating automatic suites. - Implementation sketch: Pass
PR_BASE_REFthroughci.yml, add release-branch routing, validate environment wiring and the deployment guard, and document full CI and Miri dispatches before each release candidate. - Behavioral changes worth calling out: Release PRs now run applicable Linux, macOS, Spark, Iceberg, Delta and PyArrow jobs automatically. Documentation-only changes still skip heavy jobs. Release-branch dispatches skip site deployment.
- Suggested improvements: None at the requested severity. No introduced P1/P2 issues found within this review.
Reviewed all six changed files at 1906546606a5a411bdc5a1c270830cc45694c4c5 against ee3f23930333b92f72cb0686fd97f59c639190cc. The PR is not a draft. Snapshot and live discussion checks found no existing reviews, comments or threads. Routed skill: review-comet-pr. No sibling skill applies to this CI/documentation scope.
Validation: python3 dev/ci/check-ci-config.py passed. Additional base/head policy comparisons, release-branch scenarios, CLI routing checks and disposable negative tests for the new guards passed. git diff --check passed. The checkout remains unchanged.
Exact-head CI: Comet CI passed preflight and change detection, including workflow lint and Markdown checks. The expected release-branch suites started. The label workflow, PR title check and automatic labeling passed. No failed or cancelled checks were reported.
Validation limits: Heavy CI suites remain in progress, so their final outcomes are unverified. Full Spark/native builds and runtime suites were not run locally.
Backport of #6218 to
branch-1.1.Cherry-picked from
2d1ee49a037a1a0826371fad53f4f7670943addcwithout conflicts. Of the six files it changes, five are identical onbranch-1.1and onmainjust before #6218.ci.ymldiffers only by the "Check Iceberg write report summary" step from #6155, which is not on this branch. So every changed line is identical to upstream, and only three hunk offsets inci.ymldiffer.Which issue does this PR close?
None. Part of #5327.
Rationale for this change
branch-1.1was cut at 36ab57c before #6218 merged, so itsci.ymlhas neither the release-branch routing nor thedocsguard:branch-1.1runs only the PR tier: the Linux build and the Comet suites on Spark 4.1. The other suites wait for the merge queue or the nightly, and a release branch has neither, so here they never run. chore: [branch-1.1] change version from 1.1.0-SNAPSHOT to 1.1.0 #6236 and fix: [branch-1.1] reject a file without field ids at any depth whether or not id matching is on (#6116) #6266 merged with the PR tier alone, plus the Spark 4.1 SQL suite that fix: [branch-1.1] reject a file without field ids at any depth whether or not id matching is on (#6116) #6266 was labeled for. No Iceberg, PyArrow or macOS job has run on a pull request against this branch, and neither has any Spark profile other than 4.1.docsjob has norefs/heads/mainguard. The release process now has a step that dispatchesci.ymlon the release branch before each release candidate, and a dispatch here would build this branch's docs and publish them over the website.ci.ymlmust not be dispatched onbranch-1.1until this lands.What changes are included in this PR?
The original change, so see #6218 for the details. No adaptations were needed. In short:
compute-changes.py: a pull request whose base matchesbranch-N.Mruns every job in thepr,queueandnightlytiers, still routed by the path filters.docsstays push-only, and Spark 3.4 still needs its label.ci.yml:Detect changespasses the base branch asPR_BASE_REF, and thedocsjob also requiresgithub.ref == 'refs/heads/main', so a dispatch on this branch skips the deploy.check-ci-config.py: the release-branch policy cases, a check on the variablesDetect changespasses tocompute-changes.py, and a check that fails if thedocsguard is dropped.The contributor guide is not published from this branch, because the website is built from
mainand takes only the user guide from here. The doc changes are included anyway, so the cherry-pick stays clean and the branch's docs describe its ownci.yml.How are these changes tested?
python3 dev/ci/check-ci-config.pypasses on this branch, including ci: run every tier on release-branch pull requests, and the full suite before an RC #6218's new policy cases and its checks onPR_BASE_REFand thedocsguard.actionlintreports the same 26 findings with and without this PR. All are shellcheck info and style notes in workflows this PR does not touch, and none is inci.yml.prettier --checkpasses on the three changed Markdown files.compute-changes.py, run with this PR's changed files and the environment CI sets:branch-1.1build_linux,build_linux_full,build_linux_all_profiles,build_macos,delta_gate,pyarrow_udf,spark_3_5,spark_4_0,spark_4_1,spark_4_1_hiveand all four Iceberg versionsmainbuild_linux,build_linux_fulllabeledwithrun-spark-3.4-testsspark_3_4This PR's own run is the first real exercise of the routing, as #6218 anticipated. A pull request runs the workflow files from its merge commit, and this one changes
ci.yml, so its run should include every suite except Spark 3.4, the benchmark check and the site deploy. Most of those suites have never run against this branch, so a failure in one may come from the branch rather than from this change.